Research Synthesis Methods
○ Wiley
Preprints posted in the last 90 days, ranked by how well they match Research Synthesis Methods's content profile, based on 20 papers previously published here. The average preprint has a 0.02% match score for this journal, so anything above that is already an above-average fit.
Cumpston, M. S.; Brennan, S. E.; Ryan, R.; Thomas, J.; McKenzie, J. E.
Show abstract
Introduction Systematic review authors commonly encounter situations where the data required for meta-analysis are incompletely reported (e.g. when effect estimates are reported without a measure of precision). In this circumstance, many systematic review authors use a method other than meta-analysis (e.g. vote counting), but rarely describe those methods or the rationale for selecting them. We aimed to investigate what methods authors consider when meta-analysis of all study results is not possible, and what factors influence their decisions. Methods We interviewed 12 experienced systematic review authors, editors and methodologists, presenting four scenarios in which it was not possible to combine all results using meta-analysis. Scenarios varied in the number and size of included studies, available data, and risk of bias. Participants discussed the methods they considered to summarise, synthesise and present the results; whether they would synthesise available results; and how they would draw overall conclusions. Results Factors that informed decisions included participants' overall purpose in conducting synthesis, existing beliefs about study results and synthesis methods, trust in the available data, and the decision-making needs of end users. Participants differed in which synthesis methods to use, whether they would use multiple synthesis methods, and which studies they would analyse with each method. Conclusions We identified several synthesis methods considered when meta-analysis of all results is not possible, and factors that influence the selection of methods, neither of which are routinely reported. More complete reporting of these methods and the factors informing decisions would allow readers to better understand the decisions made.
Nepomuceno, A.; Ghosal, A.; Sandoval Lentisco, A.; Ioannidis, J. P. A.
Show abstract
Empirical conclusions can depend on the many individual choices researchers make when analyzing data. Multiverse-style analyses address this by computing results across a set of specifications rather than a single one, but it is unclear how widely they are being used and how they are being implemented. We searched the Web of Science Core Collection (May 2026) for articles citing six foundational papers on different variants of multiverse-style methods and classified each of the citing articles as implementing such methods or only discussing them. For implementations, we recorded the framework used, the number of specifications, which of four decision nodes (measurement, data processing, modeling, and estimation) were varied, and other aspects such as how results were visualized and interpreted. Of the 1545 classifiable articles, 613 (39.7%) implemented a multiverse-style analysis (primarily multiverse n = 336, specification curve n = 175, vibration of effects n = 20, multimodel n = 59, and multi-analyst/many-analyst n = 23). Uptake spanned many disciplines, most often psychology (39.8%), the social sciences (20.7%), and medicine (19.2%). The number of specifications ranged from fewer than ten to more than ten thousand (median = 144, IQR 24-1248). Modeling (75%) and data-processing (62%) choices were included most often. Interpretation was predominantly descriptive, whereas formal inference (10.5%), preregistration (8.6%), and explicit attention to the defensibility of specifications (3.9%) were uncommon. Multiverse-style analyses are increasingly becoming established across the quantitative sciences, but some implementation practices can be strengthened so that these analyses become more fully transparent and more genuinely informative about the robustness of research findings.
Pan, W.; Lu, Z.; Jiang, W.; Lim, J.; Xu, L.; Wang, X.
Show abstract
In meta-analyses of continuous outcomes, the sample mean and standard deviation (SD) are essential for synthesizing effect sizes across studies. However, clinical studies frequently report alternative summary statistics, such as the median, quartiles, and range. To enable inclusion of such studies, various methods have been proposed to estimate the sample mean and SD from these reported summaries. We propose the Bayesian Order Statistics-based Estimator (BOSE), which leverages the joint likelihood of observed order statistics together with weakly informative priors to obtain the full posterior distribution for the mean and SD without relying on computationally intensive iterative procedures such as Markov chain Monte Carlo algorithms. Our numerical studies demonstrate that BOSE performs competitively with existing approaches in estimating the mean, while achieving superior performance for estimating the SD across all evaluated scenarios, particularly in small-sample settings. Under non-normal distributions including skewed, heavy-tailed, and bimodal settings with mild or moderate deviations from normality, BOSE remains robust and stable, whereas methods specifically designed for skewed distributions may become unstable or even inapplicable. Beyond point estimation, BOSE naturally provides empirically validated posterior credible intervals, enabling researchers to formally quantify uncertainty for study-level estimates and make reliable, evidence-based decisions in meta-analytic research synthesis. A publicly accessible web application implementing BOSE and competing methods is also provided to facilitate practical use in meta-analytic research.
Jiang, L.; Ying, X.; Brown, A. W.; Lan, M.; Song, W.; Menke, J.; Vorland, C.; Mayo-Wilson, E.; Kilicoglu, H.
Show abstract
Randomized controlled trials (RCTs) play a central role in assessing the benefits and harms of interventions. Incomplete reporting in RCT publications can compromise the verifiability and usefulness of RCTs. SPIRIT and CONSORT reporting guidelines aim to improve the completeness of RCT protocols and results publications, respectively. However, many RCTs are not reported completely. Checking manuscripts automatically could help authors improve the completeness of reports prior to publication. We previously annotated SPIRIT-CONSORT-TM, a corpus of 200 articles (comprising 100 protocol-results publication pairs) using 83 checklist items drawn from SPIRIT 2013 and CONSORT 2010. We also trained machine learning models to automatically assess reporting at the item level. Each checklist item can include multiple constituent elements (i.e., specific details required for that item), and an item might be considered fully reported when all of its elements are present. However, prior work does not explicitly capture or evaluate reporting at the element level. To address this gap, we extended SPIRIT-CONSORT-TM by incorporating element-level annotations and using them to assess reporting completeness (SPIRIT-CONSORT-ELM). We formulated element-level assessment as a machine reading comprehension task, operationalized through 119 questions, where each question targets a specific reporting element within a checklist item. Using the 200 articles included in SPIRIT-CONSORT-TM, two annotators independently answered 119 questions for 50 articles (25 protocol-results pairs) and resolved any discrepancies through discussion; the remaining 150 articles (75 protocol-results pairs) were assessed by a single annotator. We then developed an automated pipeline for element-level assessment using SPIRIT-CONSORT-ELM. The pipeline first applies a PubMedBERT-based model to identify sentences containing item-level reporting information, then it uses a generative large language model (LLM; GPT-5) with chain-of-thought reasoning to answer element-level questions based on the retrieved evidence. Agreement between the two annotators was high (Gwet's AC1: 0.782) and our pipeline achieved high accuracy in identifying element-level reporting evidence (F1: 0.822, Gwet's AC1: 0.796). Ablation studies indicate that chain-of-thought reasoning and the inclusion of illustrative in-context examples modestly improve LLM performance on the machine reading comprehension task. SPIRIT-CONSORT-ELM provides a benchmark for evaluating reporting guideline completeness at the element level, enabling assessment of RCT transparency beyond the simple presence or absence of checklist items and is publicly available at https://osf.io/kznx4/. The automated pipeline establishes a robust baseline for assessing RCT reporting and demonstrates potential as a practical aid for authors, reviewers, and editors to identify and address gaps in completeness and transparency of RCT reports.
Dobin, D.; Witmer, A. M.; Sweeney, F.; Ryan, T.; Cimino, A.; Haroz, E. E.; Nestadt, P. S.; Wilcox, H. C.
Show abstract
Importance. Systematic reviews and meta-analyses inform suicide-prevention policy and practice, but broad database searches are difficult to screen manually. This limits capture of upstream interventions, such as economic policies, with indirect effects on suicide. Reliable automated screening could make broader and more comprehensive evidence syntheses feasible. Objective. To develop and validate ScreenAgent, a large language model (LLM) agent for title and abstract screening, and a review-specific method for prospectively estimating screening performance. Design, Setting, and Participants. ScreenAgent was validated internally on a prospective meta-analysis, and externally on two published systematic reviews. The correct include and exclude decisions followed standard systematic-review screening methodology. Exposures. ScreenAgent, an LLM agent returning structured include-or-exclude decisions. Records it marked for inclusion were re-checked by a second, cascade pass using a higher-effort LLM. For the external reviews, the agent's prompt was tuned automatically on a small set of labeled examples. Main Outcomes and Measures. We calculated sensitivity, specificity, workload reduction (the percentage of records removed from human review), and agent-versus-human reliability via Cohen kappa. Sensitivity was estimated by direct comparison (internal) and 5-fold cross-validation (external). Results. In the internal validation, ScreenAgent identified 43 of 44 eligible studies (sensitivity 97.7%; 95% CI, 88.2%-99.6%) with a generic prompt applied without any review-specific optimization, specificity 98.0%, and a measured full-corpus workload reduction of 99.4%. The cost was $855.91 for the full 201,064-record corpus (0.43 US cents per record). Agent-versus-human-consensus agreement exceeded human-versus-human agreement (Cohen kappa 0.75 vs 0.64; percent agreement 97.3% vs 95.4%). For two external validation studies, automatic tuning resulted in a cross-validated sensitivity of 95.9% (95% CI, 90.0%-98.4%) and 97.4% (90.9%-99.3%), with workload reductions of 97.4% and 98.4%. Conclusions and Relevance. Suicide prevention efforts often require rapid consolidation of evidence because of the inherent challenges of single studies trying to prevent rare outcomes. On both internal and external validation sets, ScreenAgent identified nearly all eligible studies with human-level reliability for a fraction of a US cent per record while keeping human reviewers as the final arbiters. By making broad searches feasible and screening performance measurable beforehand, this approach can serve as a transparent methodology to strengthen the speed at which we can inform and advance suicide prevention efforts.
Otte, W. M.
Show abstract
Meta-analysis usually reduces each study to an effect estimate with a standard error and pools these by inverse-variance weighting: fixed effect (FE), random effects (RE), or unrestricted weighted least squares (UWLS). We propose information-geometric meta-integration (IGMI), representing each study by its sampling distribution, the Gaussian N(theta_i, Sigma_i), and pooling studies as a weighted Frechet mean (barycenter) under Bures-Wasserstein (BW), Fisher-Rao, or Wasserstein-Fisher-Rao (WFR) geometry. In the scalar fixed-variance case the BW barycenter mean is exactly the FE estimate; the minimized Frechet functional reproduces the Higgins-Thompson I^2 and DerSimonian-Laird tau^2 heterogeneity statistics; and a Frechet-scatter pivot reproduces the Hartung-Knapp-Sidik-Jonkman interval at m = 1 and yields an exact Hotelling F(m, K-m) region for m outcomes under proportional total covariances. WFR adds a robust outlier-resistant pool: as its length scale delta grows without bound it converges monotonically to BW, whereas finite delta gives a redescending M-estimator with rejection point exactly pi*delta. Simulations show calibrated multivariate coverage at small K, where Wald intervals undercover, and strong resistance of the equal-weight WFR pool to contamination. In 2,445 Cochrane meta-analyses, WFR most often wins leave-one-out predictive scoring. In 835 bivariate meta-analyses, the closed-form BW barycenter matches REML multivariate meta-analysis predictively and is exactly invariant to the unreported within-study correlation, unlike the likelihood estimate.
Reddy, S.; Heritier, A.
Show abstract
The rapid expansion of the medical artificial intelligence (AI) literature has outpaced our ability to judge how far published models have progressed towards clinical use. We investigated whether the translational maturity of a study can be estimated automatically from its abstract. Using PubMed, we assembled a corpus of 11,024 candidate articles, reduced it to 1,816 AI-related articles by heuristic filtering, and manually double-annotated a balanced sample of 524 articles across five maturity classes (internal validation, external validation, prospective evaluation, implementation or governance, and not applicable). Abstracts were represented as TF-IDF features and classified using multinomial logistic regression with a Lasso penalty, chosen for interpretability and suitability for a small, imbalanced dataset. On a stratified held-out test set (n = 104), the model achieved 69.2% accuracy, Cohen's kappa of 0.495, macro-F1 of 0.458 and a weighted AUC of 0.820. Performance was strong for the frequent classes but poor for the rare implementation or governance class, which the model failed to recover. A balanced manual verification of 200 large-corpus predictions confirmed this pattern, with per-class precision ranging from 82.5% (internal validation) to 5.0% (implementation or governance). An interpretable, low-resource classifier can support literature mapping but requires human oversight for advanced maturity levels.
LIn, H.; Lyu, J.
Show abstract
BackgroundQuality Control Circle (QCC) reports are often reviewed qualitatively, but reviewer workload and inter-rater variability make large-scale assessment difficult. We evaluated whether multiple large language models (LLMs) could score QCC methodological quality reliably on a designed-anchor benchmark. ObjectiveTo estimate inter-model reliability for QCC quality scoring and to assess whether model scores align with designed synthetic anchors and remain descriptively comparable to a small set of public PMC QCC reports. MethodsWe evaluated 30 synthetic QCC reports and 8 public PMC QCC reports across four primary evaluators (GPT, Gemini, Grok, DeepSeek) and one sensitivity evaluator (Claude); Claude was excluded from the primary panel because it shared the model family used during prompt development. Each synthetic case was scored across eight QCC quality dimensions in three runs per evaluator. We summarized each evaluator by median scores, then estimated ICC(A,1) across the primary panel. We also examined score-based calibration against designed anchors, keyword-assisted defect mention, leave-one-out and k=5 sensitivity, and a descriptive synthetic-versus-PMC distributional plausibility check. ResultsInter-model reliability on the primary k=4 panel was excellent: ICC(A,1) = 0.953 (95% CI 0.944 to 0.962) with 237 pooled case-dimension rows. The pre-specified k=5 sensitivity analysis including Claude was 0.954, and leave-one-out estimates within the primary panel ranged from 0.950 to 0.959. Score-based calibration against designed anchors met the prespecified target in 57/58 trap-affected case-dimension rows (98.3%). Keyword-assisted defect mention was present in 51/58 trap instances (87.9%). The synthetic-versus-PMC comparison was descriptively similar across all eight dimensions, and all dimensions met the predefined descriptive margin check. ConclusionsIn this designed-anchor pilot, multi-model LLM scoring of QCC methodological quality showed high inter-model reliability and stable alignment with synthetic anchor scores. These findings support benchmark feasibility, but they do not establish expert validity, clinical validity, or operational deployment readiness.
Elson, R.; McIntyre, K. M.; Hardingham, M. B.; Luechtefeld, T.; Lake, I. R.
Show abstract
Abstract Climate change is altering environmental conditions that influence foodborne disease transmission, yet traditional systematic reviews cannot keep pace with expanding evidence. We assessed whether an LLM-assisted workflow could generate a rapid, repeatable, and policy-relevant living evidence base for climate-sensitive foodborne disease. We combined structured PubMed searches (2010-2023), gold-standard human labelling, and iterative refinement of a GPT?4?Turbo?based auto-labeller within the SysRev platform. Pathogens of public-health importance in England were selected a priori. Model performance was evaluated against human reviewers using recall, precision, specificity, accuracy, and balanced accuracy. The refined inclusion model achieved 89{middle dot}2% recall, 59{middle dot}2% precision, 84{middle dot}5% specificity, and 85{middle dot}4% accuracy across 1,044 screened abstracts, identifying 436 studies for inclusion. Post-hoc re-evaluation of discordant abstracts showed that records excluded by the model but included during initial human screening did not meet the refined inclusion criteria. Frequently identified climate exposures included rainfall, temperature, seasonality, and humidity; norovirus, Salmonella, Campylobacter, and Cryptosporidium were the most common pathogens. An LLM-assisted workflow can generate living evidence for climate-sensitive foodborne disease with high recall and improved screening consistency. The approach is scalable, auditable, and suitable for secure institutional environments, supporting horizon scanning and climate-health risk assessment.
Sidebotham, D.; Barlow, J.
Show abstract
Background Multicentre trials in anaesthesia and critical care report low rates of statistically significant differences. This finding may partly reflect conventional sample size methods, which assume a fixed treatment effect. Assurance methods use a design prior to represent uncertainty in the expected treatment effect, which may provide a more realistic way of estimating sample sizes. Methods We calculated power curves across a range of effect sizes, design priors, and sample sizes using frequentist and Bayesian assurance methods and compared the sample sizes required to achieve 80% and 90% power to the conventional method. We standardised the design priors across effect sizes using the coefficient of variation. We derived a theoretical limit for achievable power. We validated a normal approximation to the Bayesian posterior distribution. Results Frequentist and Bayesian assurance methods produced similar power curves across all scenarios. At a coefficient of variation of 0.5 - reflecting realistic prior uncertainty in the expected effect size - both methods required sample sizes that were approximately 1.5 to 3.5 times larger than the conventional method. The theoretical power limit depends only on the coefficient of variation of the design prior and holds true across all effect sizes. The normal approximation to the Bayesian posterior distribution matched the results obtained from Markov chain Monte Carlo sampling. Conclusions Incorporating clinical uncertainty in the expected effect size substantially increases the sample size required to achieve adequate power, which has important implications for the feasibility of randomised trials in anaesthesia and critical care.
Cummins, J.; Drysdale, H.; Elson, M.; Hussey, I.; Goldacre, B.; DeVito, N. J.
Show abstract
Objective To evaluate the accuracy and cost of RegCheck, an automated large language model (LLM)-based workflow, for identifying clinical trial outcomes and detecting outcome misreporting by comparing its outputs with manual assessments from the COMPare Trials project. Design Validation study. Setting Sixty-two clinical trials originally assessed in the COMPare Trials project, sampled from five high impact general medical journals. Participants Published clinical trial reports and their corresponding prespecified registrations and/or protocols. Main outcome measures Four prespecified research questions were examined. RQ1 assessed outcome extraction recall relative to COMPare. RQ2 assessed accuracy of outcome classification as primary, secondary, or non-prespecified. RQ3 assessed accuracy of misreporting detection relative to COMPare, with additional manual adjudication of discrepancies between RegCheck and COMPare. RQ4 assessed the average per-paper cost of running the automated workflow. Results Across the validation papers, RegCheck achieved 91.2% outcome extraction recall relative to COMPare, and 83.6% outcome classification accuracy. For detection of outcome misreporting, RegCheck's overall accuracy was 85.6%. However, after resolving discrepancies with the original human judgements (which frequently favoured RegCheck's judgement), revised accuracy for outcome misreporting detection was 94.8%. The mean cost of running the workflow was 5.94 USD per paper. Conclusions RegCheck achieved high overall performance with a rigorous manual benchmark for identifying prespecified and reported outcomes in clinical trials, and detecting outcome misreporting, while operating at very low marginal cost. Adjudication of discrepant judgements suggested that RegCheck frequently identified valid issues not captured in the reference standard. Automated outcome checking may offer a scalable way to support editors, peer reviewers, and authors in detecting outcome switching and improving trial reporting.
Jaber, A.; Hughes, L.; Cameron, A. C.; Quinn, T. J.
Show abstract
Background: Systematic reviews of clinical prediction models increasingly include studies using artificial intelligence (AI) and machine learning (ML) methods alongside traditional multivariable regression approaches. A previously published Excel tool enabled standardised data extraction using the CHARMS checklist and risk of bias assessment using PROBAST. The recent publication of the PROBAST+AI framework, which distinguishes the assessment of model development quality from the assessment of model evaluation risk of bias and assesses applicability in both parts, necessitates an updated digital instrument applicable across prediction modelling methods. Methods: We updated an open-access Excel tool to incorporate the full PROBAST+AI framework. The updated template incorporates structural separation between assessment of model development quality and model evaluation risk of bias, with applicability assessed in both parts. It also incorporates updated signalling questions, including those addressing methodological issues particularly relevant to AI/ML, and automates the generation of summary tables and graphical displays. Results: The updated tool (CHARMS & PROBAST+AI Template) contains 11 worksheets and supports data extraction and appraisal for up to 30 prediction models. Dedicated, linked worksheets enable separate assessment of model development and model evaluation, with Domain 4 distinguishing among Apparent, Internal, and External evaluation settings. Key updates include dedicated assessments for predictor pre-processing, class imbalance handling and recalibration, data leakage prevention, and replication of the full model development pipeline within resampling procedures. Automated sheets dynamically format tables and summary charts covering PROBAST+AI parts. Conclusions: The CHARMS & PROBAST+AI Excel template provides a standardised, user-friendly, and rigorous digital framework for systematic reviewers appraising traditional statistical and AI-driven clinical prediction models.
Lecot, P.; Tonoli-Catez, H.; Noseda, A.; Buisse, T.; Chanut, S.; Chanel, I.
Show abstract
The exhaustive identification of evidence is central to systematic reviews, but the screening of titles and abstracts remains particularly labor intensive. Priority screening, an active learning approach that ranks records by estimated relevance, has emerged as an effective strategy to reduce screening workload. Its efficiency is commonly quantified using work saved over sampling at 100% recall (WSS@100%), representing the percentage reduction in effort compared with random screening. Although modern priority-screening models achieve high efficiency on many benchmark datasets, some reviews still exhibit low WSS@100%, indicating suboptimal retrieval. Our study sought to improve the retrieval of all relevant articles in challenging datasets to ensure better generalization of priority screening. We first showed using SYNERGY benchmark datasets that while the most advanced ELAS_h3 priority screening model from state-of-the-art ASReview LAB v.2 open-source software, efficiently retrieved most relevant articles, it struggled with the rare, final ones in challenging datasets. To address this, we tested a hybrid approach entitled SimEd AI: using ELAS_h3 for early retrieval and then applying supervised fine-tuning to the biomedical transformer BioMed-RoBERTa-base with these relevant articles to enhance the detection of the remaining difficult cases. We found that fine-tuning BioMed-RoBERTa-base model with 10 late-identified relevant and 10 hard irrelevant study titles and abstracts, enabled faster retrieval of articles of interest compared to ELAS_h3 alone. This approach increased WSS@100% from 46.5% (SD0.0%) to 83.3% (SD0.4%), while adding only an average of 22 minutes of computational time for fine-tuning and inference. The SimEd AI priority screening pipeline could be valuable for situations requiring highest possible recall. It could be particularly useful in scoping reviews with broad or diverse topics where traditional priority screening methods may miss subtle relevance signals. Further work should define a data-driven stopping rule for ending screening once the fine-tuned domain-specific transformer is applied at the final stage and assess generalizability across additional challenging datasets.
Samaan, S.; Devi, J.; Vincent, M.; Coombs, S.; Sehgal, P.; Mouhamed, M.; Rai, V.; Johnson, A. M.; Yarur, A. J.; Barnes, E. L.; Deepak, P.
Show abstract
Background: Large language models (LLMs) offer promise for systematic review data extraction, but performance in complex multidisciplinary domains and utility for clinical statement generation remain insufficiently described. Objectives: To evaluate Google NotebookLM for AI-assisted data extraction and RAND/UCLA consensus statement generation in a systematic review of IBD, obesity, and cardiometabolic comorbidities. Methods: Studies were organized into domain-specific notebooks; structured prompts generated standardized evidence tables. Two independent reviewers validated outputs against full-text articles using a four-category error classification. Cell-level accuracy and critical accuracy (cells free of major factual errors) were the primary metrics; workflow time was compared against a published conventional extraction benchmark. Concordance between AI-generated and expert-finalized statements was assessed. Results: Across 57 articles, 1,710 data cells were extracted; 151 (8.83%) were flagged, yielding 91.17% cell-level accuracy. Major factual errors occurred in only 4 cells (0.23%), for a critical accuracy of 99.77%. Most errors were minor omissions (59.6%) or incomplete extractions (30.5%); domain error rates ranged from 7.08% to 11.33%. The pipeline required 17.7 versus a projected 165.1 person-hours (89.3% reduction). PICO-structured prompting generated 70 candidate statements; 58 of 112 finalized panel statements (51.8%) were AI-derived, and 85.7% were retained in the finalized set. Conclusion: Google NotebookLM demonstrates feasibility as a primary extraction and synthesis tool in a multidisciplinary systematic review, with extractive incompleteness as the principal limitation and substantial time savings over conventional approaches. Its novel application to RAND/UCLA consensus statement generation extends AI-assisted evidence synthesis to clinical consensus generation workflow.
Clapham, M.; Oldmeadow, C.; Deeming, S.; Riveros, C.
Show abstract
Classification of research articles according to translational research stages enables funding bodies, academic and medical institutes, and policymakers to objectively assess the distribution of resources across the research spectrum. We aim to utilise Large Language Models (LLM) to classify medical research papers into translational research levels based on their titles and abstracts, comparing performance across a range of LLMs, multiple runs and a bag of words (BoW) baseline. We quantify the performance of open-weight LLMs against a human-labelled data set of 318 medical research papers. Using a description of translational levels, the LLMs showed good performance with an F1 score of 0.83 ahead of a baseline BoW approach of 0.68. We show that LLMs can accurately classify titles and abstracts into translational levels within a fully automated pipeline.
Huang, H.; Zheng, Q.; Qiu, P.; Zhao, W.; Zhang, Y.; Xie, W.; Wang, Y.; Zhang, X.; Wu, C.
Show abstract
Medical systematic reviews are central to evidence-based medicine, but they remain slow, labor-intensive, and difficult to maintain under the full Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) workflow. Recent LLM-based deep research agents offer a promising route to addressing this challenge, yet reliable deployment in medical systematic reviews remains limited by insufficient clinical domain knowledge and inconsistent adherence to evidence-based methodological standards across the full workflow. We address these gaps with MedSR-Copilot, a PRISMA-aligned multi-agent copilot that decomposes review automation into literature retrieval, coarse-to-fine screening, data extraction, Risk-of-Bias assessment, and evidence synthesis, while preserving structured intermediate artifacts throughout the workflow. We further introduce MedSR-Bench, an end-to-end benchmark for evaluating systems beyond isolated subtasks, from review input to final evidence-synthesis conclusions. MedSR-Copilot completes medical systematic reviews end-to-end under the full PRISMA workflow, achieving 63.6% human-aligned conclusions, 18.3 percentage points above the best baseline among strong general-purpose LLMs and prior automated review systems. In a human-AI collaboration study involving 23 analysis groups across four systematic review topics, MedSR-Copilot, used as a copilot, reduces end-to-end review time by 64.9% and improves final conclusion accuracy by 27.4 percentage points compared with routine-practice workflows. Together, these results demonstrate the reliability and efficiency of MedSR-Copilot as a medical research copilot and suggest a practical path toward trustworthy review automation.
Gorici, M.; Torres Espin, A.; Oremus, M.
Show abstract
Introduction. Systematic reviews are essential for informing health policy and practice. Artificial intelligence (AI) automates the article screening process and produces time savings, although the performance of AI screening compared to traditional human screening remains uncertain. We undertook this study to compare the performance of two agentic AI tools, namely Loon LensTM and Catchii, to one another and to humans at the title and abstract screening level. We also compared Loon Lens to humans at the full-text screening level. Methods. We developed a de novo research question on the association between any of three ambient air pollutants (carbon monoxide, ozone, nitrogen dioxide) and the onset or worsening of Parkinson disease. A health sciences librarian developed the literature search strategy and we proceded with human screening guided by PRISMA. We uploaded the retrieved citations and the eligibility criteria to both AI tools and compared screening results using sensitivity, specificity, positive predictive value, negative predictive value, concordance, kappa, and F1 score. We compared the calculated performance statistics to those obtained by naive guessing and regressed concordance (agree or disagree with the human reference standard) onto confidence scores provided by Loon Lens, which assigned a confidence level (Very High, High, Medium, or Low) to each of its screening decisions. Human screening was the reference standard against both AI tools and Catchii was the reference standard against Loon Lens. Results. At title and abstract screening, Loon Lens outperformed Catchii when humans were the reference standard. At full-text screening, most disagreements centered around articles Loon Lens included and humans excluded. At both screening levels, higher confidence scores were associated with lower odds of disagreement between Loon Lens and human screeners. Discussion. Given the panoply of available AI screening tools and their differential performance, plus the rapidly evolving nature of AI technology, researchers should pilot test their chosen tool at the start of each review. Sensitivity, kappa, and F1 are the optimal performance statistics to employ, especially at title and abstract screening, where the imbalance between proportions of included and excluded citations can inflate concordance and negative predictive value.
Li, S.; Zhang, W.; Xing, X.; Shen, Z.; Wang, Y.; Chen, Z.; Neto, O.; Yu, Y.; Wu, C.; Lin, L.
Show abstract
Background Late-stage cancer incidence is being considered as an earlier endpoint in cancer-screening trials, but its trial-level association with cancer-specific mortality may depend on evidence selection and endpoint harmonization. We evaluated the robustness of this association to source-verified additions. Methods We reconstructed the PubMed corpus underlying a 41-comparison review. Gemini 3.1 Pro Preview was used only to prioritize reports for blinded human reassessment. Reviewers determined eligibility, linked reports from the same trial, harmonized endpoints, and verified comparison-level data. We recalculated unweighted Pearson correlations overall and by cancer type after adding earliest-compatible trial comparisons. Results Among 1209 candidate records, 996 PDFs were assessed. Thirty-three reports absent from the source review were prioritized; 26 were eligible, representing 18 trials, and 8 provided compatible comparisons. Adding these comparisons increased the dataset from 41 to 49 and attenuated the overall correlation from 0.73 (95% confidence interval [CI] = 0.55 to 0.85) to 0.59 (95% CI = 0.37 to 0.75). Updated correlations were 0.49 (95% CI = -0.26 to 0.87) for breast, -0.23 (95% CI = -0.71 to 0.40) for colorectal, and 0.83 (95% CI = 0.54 to 0.95) for lung cancer. One sparse-event comparison influenced the colorectal estimate. Conclusions The overall association was sensitive to evidence composition, and cancer-specific stability varied. Late-stage incidence should be evaluated by cancer type and with prespecified sensitivity analyses for evidence selection and endpoint definitions. Model-assisted prioritization cannot replace human eligibility review, trial reconciliation, and source verification.
Malec, S. A.; Pradhan, M.; Upadhayaya, R.; Metzger, V.
Show abstract
Objective: Observational studies are essential for investigating risk factors for Alzheimer's disease and related dementias (ADRD), but inconsistent reporting and selection of covariates can contribute to residual confounding, omitted-variable bias, and reduced reproducibility. We developed and evaluated VAREX (Variable Extraction), a large language model (LLM)-based information extraction framework designed to automatically identify exposures, outcomes, and covariates from epidemiologic studies and populate structured evidence repositories. Materials and Methods: VAREX combines retrieval-augmented generation, biomedical language-model embeddings, semantic chunking, cross-encoder reranking, and prompt-engineered LLM workflows to extract epidemiologic variables from full-text biomedical articles. The framework was evaluated using a reference-standard corpus of observational studies examining blood pressure variability (BPV) and Alzheimer's disease-related dementias (ADRD), together with external validation datasets involving other exposure-outcome relationships. Extracted variables were compared with independently curated human reference standards using semantic matching and one-to-one assignment procedures. Covariates were additionally classified into ten epidemiologically relevant semantic categories. Results: In the primary BPV[->]ADRD corpus (10 studies), VAREX achieved a precision of 0.91, recall of 0.84, and F1-score of 0.87 for variable extraction. Covariate classification accuracy was 0.90, yielding a strict extraction-and-classification F1-score of 0.78. External validation datasets demonstrated comparable performance across diverse epidemiologic domains, with extraction F1-scores ranging from 0.73 to 0.85. Category-level performance was strongest for health behaviors (F1=0.96), sociodemographic variables (F1=0.90), and medication exposures (F1=0.89). Compared with published estimates of manual systematic-review effort, VAREX reduced processing time from approximately 61 minutes to 9 minutes per article, representing an 85.7% reduction in review time. Discussion: These findings demonstrate that LLM-based information extraction can accurately identify and classify epidemiologic variables across heterogeneous observational-study designs. Automated extraction enables scalable construction of structured repositories of exposures, outcomes, and covariates while substantially reducing the labor required for evidence synthesis and systematic reviews. Conclusion: VAREX provides an effective framework for automated extraction and classification of epidemiologic variables from the biomedical literature. By supporting large-scale evidence synthesis and structured knowledge resource development, VAREX may facilitate more rigorous observational research, improved confounder identification, and enhanced reproducibility in epidemiology.
Oparah, C.; O'Keefe, H.; Agbeleye, O.; Nesworthy, J.; Norman, G.; Kunonga, T. P.
Show abstract
Clinical trials often enrol populations that differ from those who ultimately receive the interventions, raising concerns about external validity and health equity. Trial registries could provide an early opportunity to assess representativeness, but it is unclear whether registry data contain sufficient information to enable such assessments. This study evaluated the feasibility of using registry data to assess representativeness in Phase II and III pharmacological randomised controlled trials. A search of ClinicalTrials.gov from December 2024 to January 2025 identified trials with results posted after 1 January 2023 across cardiovascular disease (CVD) excluding stroke, diabetes mellitus, and selected mental health disorders. Of 1,328 records screened, 98 trials met inclusion criteria (51 Phase III, 47 Phase II). Reporting completeness was variable, particularly in Phase II studies. CVD and diabetes trials predominantly included middle-aged to older adults, while mental health trials recruited mainly individuals aged 36 to 50 years. Across CVD and mental health trials, participants were largely male. Reporting of BMI, contraception, and comorbidity criteria was inconsistent, though available data suggested these factors influenced sample composition. Fewer than 10% of trials reported equity-relevant characteristics beyond age and sex, and none addressed intersectionality. Assessing equity using registry data is feasible but constrained by incomplete and inconsistent reporting.